prompt-cache-bench Longitudinal benchmark registry for LLM inference systems

Open Longitudinal Benchmark Registry

Tracking prompt caching and provider routing across models, providers, and time.

prompt-cache-bench is designed as a long-running, reproducible registry for LLM inference systems. It will accumulate hundreds of model/provider A/B comparisons, raw datasets, report pages, figures, and runnable experiment code under a consistent methodology.

Experiment Registry

Each row represents a versioned benchmark artifact. Future experiments should be appended here with stable report URLs, raw datasets, reproducible scripts, and manifest checksums.

Status Model A/B Pair Design Reports Data
Published google/gemini-3-flash-preview Infron vs OpenRouter 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published qwen/qwen3.6-flash Infron vs OpenRouter 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published qwen/qwen3.6-35b-a3b Infron vs OpenRouter 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published google/gemma-4-26b-a4b Infron vs OpenRouter 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only; OpenRouter alias google/gemma-4-26b-a4b-it. EN HTML ZH HTML EN MD ZH MD Reports Dataset
Published minimax/minimax-m3 Infron vs OpenRouter 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published minimax/minimax-m2.7 Infron vs OpenRouter 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published minimax/minimax-m2.5 Infron vs OpenRouter 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published xiaomi/mimo-v2.5 Infron vs OpenRouter 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published moonshotai/kimi-k2.7-code Infron vs OpenRouter 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. EN HTML ZH HTML EN MD ZH MD Reports Dataset
Published moonshotai/kimi-k2.6 Infron vs OpenRouter 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. EN HTML ZH HTML EN MD ZH MD Reports Dataset
Published moonshotai/kimi-k2.5 Infron vs OpenRouter 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published z-ai/glm-4.7 Infron vs OpenRouter 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published z-ai/glm-5.1 Infron vs OpenRouter 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published z-ai/glm-5 Infron vs OpenRouter 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published qwen/qwen3.5-27b Infron vs OpenRouter 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. EN HTML ZH HTML EN MD ZH MD Reports Dataset
Published openai/gpt-5.4-nano Infron vs OpenRouter 4x50 paired streaming Chat Completions; routing sort modes; default reasoning/thinking; prompt-length tiers; /v1/chat/completions only. EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published openai/gpt-4o-mini Infron vs OpenRouter 4 x 50, streaming, routing sort: throughput / price / latency / ttft, platform-default reasoning, short / medium / long prompt tiers, Chat Completions only EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published qwen/qwen3-next-80b-a3b-instruct Infron vs OpenRouter 4 x 50, streaming, routing sort: throughput / price / latency / ttft, platform-default reasoning, short / medium / long prompt tiers, Chat Completions only EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published qwen/qwen3.5-plus Infron vs OpenRouter 4 x 50, streaming, routing sort: throughput / price / latency / ttft, platform-default reasoning, short / medium / long prompt tiers, Chat Completions only EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published openai/gpt-5.4-mini Infron vs OpenRouter 4 x 50, streaming, routing sort: throughput / price / latency / ttft, platform-default reasoning, short / medium / long prompt tiers, Chat Completions only EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published deepseek/deepseek-v4-pro Infron vs OpenRouter 4 x 50, streaming, routing sort: throughput / price / latency / ttft, platform-default reasoning, short / medium / long prompt tiers, Chat Completions only EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published z-ai/glm-5.2 Infron vs OpenRouter 4 x 50, streaming, routing sort: throughput / price / latency / ttft, platform-default reasoning, short / medium / long prompt tiers EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Published deepseek/deepseek-v4-flash Infron vs OpenRouter 4 x 50, streaming, routing sort: throughput / price / latency / ttft, platform-default reasoning, short / medium / long prompt tiers, Chat Completions only EN HTML · ZH HTML · EN MD · ZH MD · Reports Dataset
Not published llama/*, claude/*, gpt/*, other qwen/* Provider matrix expansion Matched-payload A/B, streaming TTFT, provider attribution, cost breakdown To be published To be published

Coverage Model

The homepage is organized around the registry dimensions that will scale as the benchmark grows.

Model FamiliesMulti-vendorOne directory per provider/model id, preserving canonical model naming.
A/B PairsVersionedEach comparison has its own report, data, figures, code snapshot, and manifest.
MetricsComparableStrict input-token controls keep cache, cost, throughput, latency, and TTFT comparable.
Time SeriesRepeatableRepeated runs can track provider drift, cache stability, and routing changes over time.

Infron Papers

Research papers from Infron related to LLM gateway evaluation, routing, and inference systems.

Paper Authors Published Subjects Links
SEAR: Schema-Based Evaluation and Routing for LLM Gateways Zecheng Zhang, Han Zheng, Yue Xu 2026-03-20 cs.DB, cs.AI, cs.CL arXiv · PDF

Reproducibility Protocol

The project treats every benchmark as a controlled research artifact, not a one-off dashboard snapshot.

1

Fixed payloads

Each routing mode uses stable payload SHA256 values and sends two identical prompts per round.

2

Strict pairing

Only `sort/group/round` pairs with equal response-side prompt tokens are retained.

3

Observable telemetry

Usage, cost, TTFT, latency, provider identifiers, and cache tokens are preserved when returned.

4

Auditable release

Reports, raw data, code snapshots, and checksums are committed together.

Expansion Roadmap

The registry is ready for broader model/provider coverage while preserving the same evidence standard.

Model matrix

Add Qwen, Llama, Claude, GPT, Gemini, Mistral, and domain-specific model families under the same directory convention.

Provider matrix

Track cross-platform comparisons and upstream provider behavior, including routing drift over repeated runs.

Longitudinal runs

Repeat experiments over time to observe cache TTL, provider fallback, price movement, latency tails, and throughput stability.

长期开放 Benchmark Registry

持续追踪模型、Provider 与时间维度下的 Prompt Caching 和路由表现。

prompt-cache-bench 是一个长期、可复现的 LLM 推理系统 benchmark registry。它会持续沉淀数百个模型与 provider 之间的 A/B 对比,统一保存原始数据、报告、图表、实验代码与校验清单。

实验注册表

每一行代表一个带版本的 benchmark 工件。未来实验会继续追加稳定报告 URL、原始数据集、复现实验代码和 manifest checksum。

状态 模型 A/B Pair 实验设计 报告 数据
已发布 google/gemini-3-flash-preview Infron vs OpenRouter 4x50 配对 streaming Chat Completions;routing sort 模式;默认 reasoning/thinking;prompt-length tiers;仅 /v1/chat/completions。 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
已发布 qwen/qwen3.6-35b-a3b Infron vs OpenRouter 4x50 配对 streaming Chat Completions;routing sort modes;默认 reasoning/thinking;prompt-length tiers;仅 /v1/chat/completions。 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
已发布 google/gemma-4-26b-a4b Infron vs OpenRouter 4x50 配对 streaming Chat Completions;routing sort modes;默认 reasoning/thinking;prompt-length tiers;仅 /v1/chat/completions;OpenRouter 使用 alias google/gemma-4-26b-a4b-it。 中文 HTML EN HTML 中文 MD EN MD 报告目录 数据集
已发布 minimax/minimax-m3 Infron vs OpenRouter 4x50 配对 streaming Chat Completions;routing sort 模式;默认 reasoning/thinking;Prompt 长度分层;仅 /v1/chat/completions。 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
已发布 minimax/minimax-m2.7 Infron vs OpenRouter 4x50 配对 streaming Chat Completions;routing sort 模式;默认 reasoning/thinking;Prompt 长度分层;仅 /v1/chat/completions。 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
已发布 minimax/minimax-m2.5 Infron vs OpenRouter 4x50 配对 streaming Chat Completions;routing sort 模式;默认 reasoning/thinking;Prompt 长度分层;仅 /v1/chat/completions。 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
已发布 xiaomi/mimo-v2.5 Infron vs OpenRouter 4x50 配对流式 Chat Completions;routing sort 模式;默认 reasoning/thinking;prompt 长度分层;仅 /v1/chat/completions。 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
已发布 moonshotai/kimi-k2.7-code Infron vs OpenRouter 4x50 配对流式 Chat Completions;routing sort 模式;默认 reasoning/thinking;prompt 长度分层;仅 /v1/chat/completions。 中文 HTML EN HTML 中文 MD EN MD 报告目录 数据集
已发布 moonshotai/kimi-k2.6 Infron vs OpenRouter 4x50 配对 streaming Chat Completions;routing sort modes;默认 reasoning/thinking;prompt-length tiers;仅 /v1/chat/completions。 中文 HTML EN HTML 中文 MD EN MD 报告目录 数据集
已发布 moonshotai/kimi-k2.5 Infron vs OpenRouter 4x50 配对流式 Chat Completions;routing sort 模式;默认 reasoning/thinking;Prompt 长度分层;仅 /v1/chat/completions。 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
已发布 z-ai/glm-4.7 Infron vs OpenRouter 4x50 配对流式 Chat Completions;routing sort 模式;默认 reasoning/thinking;Prompt 长度分层;仅 /v1/chat/completions。 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
已发布 z-ai/glm-5.1 Infron vs OpenRouter 4x50 配对流式 Chat Completions;routing sort 模式;默认 reasoning/thinking;Prompt 长度分层;仅 /v1/chat/completions。 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
已发布 z-ai/glm-5 Infron vs OpenRouter 4x50 配对流式 Chat Completions;routing sort 模式;默认 reasoning/thinking;Prompt 长度分层;仅 /v1/chat/completions。 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
已发布 openai/gpt-5.4-nano Infron vs OpenRouter 4x50 配对流式 Chat Completions;routing sort 模式;默认 reasoning/thinking;Prompt 长度分层;仅 /v1/chat/completions。 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
已发布 openai/gpt-4o-mini Infron vs OpenRouter 4 x 50,streaming,routing sort: throughput / price / latency / ttft,平台默认 reasoning,short / medium / long prompt tiers,仅 Chat Completions 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
已发布 qwen/qwen3-next-80b-a3b-instruct Infron vs OpenRouter 4 x 50,streaming,routing sort: throughput / price / latency / ttft,平台默认 reasoning,short / medium / long prompt tiers,仅 Chat Completions 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
已发布 qwen/qwen3.5-plus Infron vs OpenRouter 4 x 50,streaming,routing sort: throughput / price / latency / ttft,平台默认 reasoning,short / medium / long prompt tiers,仅 Chat Completions 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
已发布 openai/gpt-5.4-mini Infron vs OpenRouter 4 x 50, streaming, routing sort: throughput / price / latency / ttft, platform-default reasoning, short / medium / long prompt tiers, Chat Completions only 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
已发布 deepseek/deepseek-v4-pro Infron vs OpenRouter 4 x 50,streaming,routing sort: throughput / price / latency / ttft,平台默认 reasoning,short / medium / long prompt tiers,仅 Chat Completions 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
已发布 z-ai/glm-5.2 Infron vs OpenRouter 4 x 50,streaming,routing sort: throughput / price / latency / ttft,平台默认 reasoning,short / medium / long prompt tiers 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
已发布 deepseek/deepseek-v4-flash Infron vs OpenRouter 4 x 50,streaming,routing sort: throughput / price / latency / ttft,平台默认 reasoning,short / medium / long prompt tiers,仅 Chat Completions 中文 HTML · EN HTML · 中文 MD · EN MD · 报告目录 数据集
未发布 llama/*, claude/*, gpt/*, other qwen/* Provider matrix expansion Matched-payload A/B、streaming TTFT、provider attribution、cost breakdown 待发布 待发布

覆盖模型

首页按照 registry 的长期维度组织,后续实验数量增长时仍能保持清晰结构。

Model Families多模型族每个 provider/model id 使用独立目录,保留规范模型命名。
A/B Pairs版本化每个对比都包含报告、数据、图表、代码快照和 manifest。
Metrics可比较严格 input-token 控制,使缓存、成本、吞吐、时延和 TTFT 可比。
Time Series可重复重复实验可追踪 provider 漂移、缓存稳定性和路由变化。

Infron 论文

Infron 关于 LLM gateway 评估、路由与推理系统的研究论文,后续论文会继续追加到这里。

论文 作者 发布日期 主题 链接
SEAR: Schema-Based Evaluation and Routing for LLM Gateways Zecheng Zhang, Han Zheng, Yue Xu 2026-03-20 cs.DB, cs.AI, cs.CL arXiv · PDF

可复现协议

本项目把每一次 benchmark 都组织为可审计研究工件,而不是一次性的 dashboard 截图。

1

固定 Payload

每一种 routing mode 都使用稳定 payload SHA256,每轮发送两次完全相同的 prompt。

2

严格配对

只有 `sort/group/round` 下响应侧 prompt tokens 完全一致的 A/B 样本进入统计。

3

可观测 Telemetry

保留 usage、cost、TTFT、latency、provider 标识和 cache tokens 等响应字段。

4

可审计发布

报告、原始数据、代码快照和 checksum 一起提交。

扩展路线

Registry 已经为更大规模的模型与 provider 覆盖做好结构准备,同时保持同一证据标准。

模型矩阵

继续加入 Qwen、Llama、Claude、GPT、Gemini、Mistral 以及垂直领域模型族。

Provider 矩阵

追踪跨平台对比与上游 provider 行为,包括多轮重复实验中的 routing drift。

长期时间序列

重复实验以观察 cache TTL、provider fallback、价格变化、尾延迟和吞吐稳定性。